Hands-on Genomics Tutorials
Welcome to a collection of hands-on bioinformatics exercises designed to build practical skills in genomics, data analysis, and computational biology. The sequence moves from raw sequencing reads to filtered variants and population-genomic interpretation, using realistic workflows and modern tools.
Each tutorial is written for guided or independent use. Commands can be inspected and copied directly, but the emphasis remains on understanding the biological purpose of each step, checking quality, and interpreting the output responsibly.
Pipeline · Part 1
FASTQ to VCF
Access an HPC system, inspect FASTQ files, run quality control, trim reads and map them to a reference genome.
Pipeline · Part 2
BAM to filtered variants
Filter alignments, call and assess variants, explore VCF files and perform an initial PCA.
Population genomics
Population differentiation and functional insights
Move from SNP-level differentiation to biological interpretation: estimate \(F_{ST}\), identify high-differentiation regions, connect candidates with functional evidence, and integrate climate data.
Research project · Part 1
Candidate-gene discovery
Compare nucleotide diversity (\(\pi\)), Tajima’s D and \(F_{ST}\) between defence and background genomic regions. Detect differentiation outliers, map them to annotated genes, and interpret evolutionary signals using public Arabidopsis thaliana data from the 1001 Genomes Project.
Research project · Part 2
GO enrichment of high-\(F_{ST}\) genes
Summarise SNP-level \(F_{ST}\) at gene level, test Gene Ontology enrichment with topGO’s Fisher and KS statistics, and interpret enriched biological processes in the context of defence-related selection.
RADseq workflow
De novo RADseq workflow
Integrate Stacks, RADstackshelpR and SNPfiltR for parameter optimisation, de novo SNP discovery, systematic filtering, and quality validation in non-model species.
Command line
AWK + SED quick reference
Consult the AWK and SED commands used throughout the pipeline, organised by biological purpose with reusable syntax, explanatory notes, and common pitfalls.
These materials support teaching and self-study. Computational tutorial pages with pre-rendered figures are preserved as static HTML, so publishing the website does not require access to the original HPC directories or course datasets. The interactive learning tools and course companions are maintained separately.